Papers by Ngoc Thang Vu
Exploring Segmentation Approaches for Neural Machine Translation of Code-Switched Egyptian Arabic-English Text (2023.eacl-main)
Copied to clipboard
| Challenge: | Code-switching (CS) is a problem in machine translation, but its performance is not investigated for CS settings. |
| Approach: | They propose to use morphological segmentation techniques for machine translation tasks . they compare morphology-based and frequency-based segmentation for MT tasks based on data size . |
| Outcome: | The proposed approach performs best in MT tasks but under-performs in other languages. |
ADVISER: A Dialog System Framework for Education & Research (P19-3)
Copied to clipboard
Daniel Ortega, Dirk Väth, Gianna Weber, Lindsey Vanderlyn, Maximilian Schmidt, Moritz Völkel, Zorica Karacevic, Ngoc Thang Vu
| Challenge: | In this paper, we focus on task-oriented dialog systems, although our framework allows easy integration of non-task dialog systems and their combination. |
| Approach: | They propose an open source dialog system framework for education and research that supports multi-domain task-oriented conversations in two languages. |
| Outcome: | The proposed framework supports multi-domain task-oriented conversations in two languages and is open source for education and research. |
Toward Implicit Reference in Dialog: A Survey of Methods and Data (2022.aacl-main)
Copied to clipboard
| Challenge: | In natural language, speakers often leave out information that is understood by the other party through the shared context. |
| Approach: | They propose to use omitted entities as implicit references in dialogs to improve language processing. |
| Outcome: | The proposed method is based on a set of experiments which show that the proposed method has a high level of accuracy and is a success. |
F1 is Not Enough! Models and Evaluation Towards User-Centered Explainable Question Answering (2020.emnlp-main)
Copied to clipboard
| Challenge: | Existing models and evaluation settings have shortcomings regarding the coupling of answer and explanation which might cause serious issues in user experience. |
| Approach: | They propose a hierarchical model and a new regularization term to strengthen the coupling of answer and explanation and two evaluation scores to quantify the couple. |
| Outcome: | The proposed model strengthens the answer-explanation coupling and provides evaluation scores that align with user experience. |
Beyond Accuracy: A Consolidated Tool for Visual Question Answering Benchmarking (2021.emnlp-demo)
Copied to clipboard
| Challenge: | Existing evaluation tools for general Visual Question Answering (VQA) systems are limited to answering accuracy, but they can be used to evaluate performance in real-world scenarios. |
| Approach: | They propose a browser-based benchmarking tool with an API for easy integration of new models and datasets to keep up with the fast-changing landscape of VQA. |
| Outcome: | The proposed tool tests generalization capabilities of models across multiple datasets and includes metrics that measure biases and uncertainty to further explain model behavior. |
Neighboring Words Affect Human Interpretation of Saliency Explanations (2023.findings-acl)
Copied to clipboard
| Challenge: | Recent studies found that superficial factors such as word length can distort human interpretation of the communicated saliency scores. |
| Approach: | They conduct a user study to examine how the marking of a word’s *neighboring words* affect the explainee’s perception of the word’ s importance in the context of . a saliency explanation. |
| Outcome: | The findings question whether text-based saliency explanations should continue to be communicated at word level and inform future research on alternative methods. |
Fine-tuning BERT for Low-Resource Natural Language Understanding via Active Learning (2020.coling-main)
Copied to clipboard
| Challenge: | Recent work has explored the suitability of pre-trained language models in low resource settings with less than 1,000 training data points. |
| Approach: | They propose to use pool-based active learning to speed up training while keeping the cost of labeling new data constant. |
| Outcome: | The proposed model can be fine-tuned to optimize for low-resource settings while keeping the cost of labeling constant. |
»textklang« – Towards a Multi-Modal Exploration Platform for German Poetry (2022.lrec-1)
Copied to clipboard
Nadja Schauffler, Toni Bernhart, Andre Blessing, Gunilla Eschenbach, Markus Gärtner, Kerstin Jung, Anna Kinder, Julia Koch, Sandra Richter, Gabriel Viehhauser, Ngoc Thang Vu, Lorenz Wesemann, Jonas Kuhn
| Challenge: | »textklang« aims to explore the relationship between written text and its potential and actual sonic realisation in lyric poetry . the platform will combine three modalities: the poetic text, the audio signal of a recorded recitation and, at a later stage, music scores of . musical setting of lyrical poetry. |
| Approach: | They propose to combine a multi-modal corpus of German lyric poetry from the Romantic era with a platform for systematic exploration. |
| Outcome: | The platform will combine the poetic text, the audio signal of a recorded recitation and, at a later stage, music scores of . a musical setting of lyric poetry. |
Meta Learning and Its Applications to Natural Language Processing (2021.acl-tutorials)
Copied to clipboard
| Challenge: | Meta-learning is a new technique that aims to learn better learning algorithms, including better parameter initialization, optimization strategy, network architecture, distance metrics, and beyond. |
| Approach: | This tutorial introduces Meta-learning approaches and the theory behind them, and then reviews the works of applying this technology to NLP problems. |
| Outcome: | This tutorial will introduce Meta-learning approaches and the theory behind them, and then review the works of applying this technology to NLP problems. |
Prompting-based Synthetic Data Generation for Few-Shot Question Answering (2024.lrec-main)
Copied to clipboard
| Challenge: | Language models have boosted the performance of Question Answering, but data annotation is costly. |
| Approach: | They propose to use large language models to improve Question Answering performance . they argue that domain-agnostic knowledge from LMs is sufficient to create a well-curated dataset. |
| Outcome: | The proposed model outperforms state-of-the-art approaches on few-shot Question Answering. |
A Survey of Code-switched Arabic NLP: Progress, Challenges, and Future Directions (2025.coling-main)
Copied to clipboard
| Challenge: | Code-switching (CSW) is a common linguistic phenomenon in multilingual societies . current literature on CSW in the arab world is limited to the Arabic language . |
| Approach: | They present a review of the literature in the field of code-switched Arabic NLP . they propose recommendations for future research . |
| Outcome: | This review provides a broad perspective on the current literature in the field of code-switched Arabic NLP . it also provides recommendations for future research . |
Intrinsic Subgraph Generation for Interpretable Graph Based Visual Question Answering (2024.lrec-main)
Copied to clipboard
| Challenge: | Visual Question Answering (VQA) is acknowledged as a challenging multi-modal task for Machine Learning (ML). |
| Approach: | They propose an interpretable approach for graph-based Visual Question Answering . their model is designed to intrinsically produce a subgraph during the question-answering process as its explanation . |
| Outcome: | The proposed model outperforms existing explainable methods on a graph-based VQA dataset. |
Ethical Considerations for Machine Translation of Indigenous Languages: Giving a Voice to the Speakers (2023.acl-long)
Copied to clipboard
| Challenge: | In recent years, machine translation has become very successful for high-resource language pairs. |
| Approach: | They conduct interviews with community leaders, teachers, and language activists to shed light on ethical considerations for the automatic translation of Indigenous languages. |
| Outcome: | The results show that the inclusion of native speakers and community members is vital to performing better and more ethical research on Indigenous languages. |
Introducing Two Vietnamese Datasets for Evaluating Semantic Models of (Dis-)Similarity and Relatedness (N18-2)
Copied to clipboard
| Challenge: | Existing datasets for low-resource language Vietnamese assess semantic similarity . a dataset for word pairs with similarity levels is needed to evaluate these models . |
| Approach: | They present two new datasets for the low-resource language Vietnamese to assess models of semantic similarity. |
| Outcome: | The two datasets are comparable to the English datasets. |
Low-Resource Multilingual and Zero-Shot Multispeaker TTS (2022.aacl-main)
Copied to clipboard
| Challenge: | Currently, the amount of data needed for TTS is limited to the vast majority of the spoken languages. |
| Approach: | They propose to use language agnostic meta learning procedure to learn speaking a new language with just 5 minutes of training data while retaining the ability to infer the voice of even unseen speakers. |
| Outcome: | The proposed approach is able to learn speaking a new language using just 5 minutes of training data while retaining the ability to infer the voice of even unseen speakers in the newly learned language. |
Cairo Student Code-Switch (CSCS) Corpus: An Annotated Egyptian Arabic-English Corpus (2020.lrec-1)
Copied to clipboard
| Challenge: | Code-switching is a phenomenon commonly observed in the Arabicspeaking world . there is still a huge gap in the available resources and NLP applications . |
| Approach: | They propose a corpus of Egyptian- Arabic code-switch speech data that is fully tokenized, lemmatized and annotated for part-of-speech tags. |
| Outcome: | The proposed corpus of Egyptian- Arabic code-switch speech data is fully tokenized, lemmatized and annotated for part-of-speech tags. |
Fast and Accurate Non-Projective Dependency Tree Linearization (2020.acl-main)
Copied to clipboard
| Challenge: | Existing methods for decoding dependency trees are 10 times faster than current ones. |
| Approach: | They propose a graph-based method to tackle a dependency tree linearization task . they propose to solve a Traveling Salesman Problem and combine the solution into a projective tree . |
| Outcome: | The proposed method outperforms the state-of-the-art linearizer while being 10 times faster in training and decoding. |
Few-shot Learning for Slot Tagging with Attentive Relational Network (2021.eacl-main)
Copied to clipboard
| Challenge: | Recent studies have used metric-based learning in computer vision but not slot tagging. |
| Approach: | They propose a metric-based learning architecture that extends relation networks by leveraging pretrained contextual embeddings such as ELMO and BERT and by using attention mechanism. |
| Outcome: | The proposed method outperforms state-of-the-art methods on SNIPS data on a slot tagging task with a large amount of hand-labeled data. |
ADVISER: A Toolkit for Developing Multi-modal, Multi-domain and Socially-engaged Conversational Agents (2020.acl-demos)
Copied to clipboard
Chia-Yu Li, Daniel Ortega, Dirk Väth, Florian Lux, Lindsey Vanderlyn, Maximilian Schmidt, Michael Neumann, Moritz Völkel, Pavel Denisov, Sabrina Jenne, Zorica Kacarevic, Ngoc Thang Vu
| Challenge: | Existing toolkits for developing dialog systems are limited to core components and do not support multi-modal processing and social signals. |
| Approach: | They propose to use ADVISER to develop multi-modal dialog agents using multi-text and social signals. |
| Outcome: | The proposed toolkit is flexible, easy to use, and easy to extend for linguists and cognitive scientists, thereby providing a flexible platform for collaborative research. |
Conversational Tree Search: A New Hybrid Dialog Task (2023.eacl-main)
Copied to clipboard
| Challenge: | Existing conversational interfaces are limited to FAQs and dialogs, allowing users to search for specific questions. |
| Approach: | They propose a task that bridges the gap between FAQ-style information retrieval and task-oriented dialog. |
| Outcome: | The proposed task bridges the gap between FAQ-style information retrieval and task-oriented dialog. |
Explaining Pre-Trained Language Models with Attribution Scores: An Analysis in Low-Resource Settings (2024.lrec-main)
Copied to clipboard
| Challenge: | Currently, prompt-based models are gaining popularity due to their easier adaptability in low-resource settings. |
| Approach: | They analyze attribution scores extracted from prompt-based models w.r.t. plausibility and faithfulness and compare them with attribution score extracted from fine-tuned models and large language models. |
| Outcome: | The proposed model outperforms attention and Integrated Gradients in plausibility and faithfulness, while fine-tuning models are harder to explain in low-resource settings. |
ArzEn: A Speech Corpus for Code-switched Egyptian Arabic-English (2020.lrec-1)
Copied to clipboard
| Challenge: | a corpus of Arabic-English code-switching (CS) spontaneous speech is collected in an Egyptian university soundproof room . the language in Egypt is rather complex and poses many challenges to natural language processing (NLP) |
| Approach: | They present an Egyptian Arabic-English code-switching (CS) spontaneous speech corpus. |
| Outcome: | The proposed corpus is designed to be used in automatic speech recognition systems . it provides a useful resource for analyzing the CS phenomenon from linguistic, sociological, and psychological perspectives. |
IMSurReal: IMS at the Surface Realization Shared Task 2019 (D19-63)
Copied to clipboard
| Challenge: | a system for shallow and deep completion is presented for the Surface Realization Shared Task 2019 . the system achieves state-of-the-art performance without using external data. |
| Approach: | They propose a surface realization system that takes five steps without external data . they perform detailed error analysis revealing correlation between word order freedom and difficulty . |
| Outcome: | The proposed system achieves state-of-the-art without external data . it achieves highest BLEU scores on tokenized text and human evaluation on four languages . |
Discrete Subgraph Sampling for Interpretable Graph based Visual Question Answering (2025.coling-main)
Copied to clipboard
| Challenge: | XAI aims to make machine learning models more transparent, but interpretable approaches are relatively rare. |
| Approach: | They integrate discrete subset sampling methods into a graph-based visual question answering system to evaluate their interpretability. |
| Outcome: | The proposed methods mitigate trade-off between interpretability and answer accuracy while achieving strong co-occurrences between answer and question tokens. |
It’s What You Say and How You Say It: Investigating the Effect of Linguistic vs. Behavioral Adaptation in Task-Oriented Chatbots (2025.coling-main)
Copied to clipboard
| Challenge: | linguistic adaptation is not known to have a positive impact on dialog success and user perception. |
| Approach: | They evaluate subjective and objective aspects of dialog success and user perceptions through a user study . they also examine linguistic adaptations of dialog agents to determine which aspects influence user perception . |
| Outcome: | The proposed agents can differ in their level of formality and their linguistic style. |
DIAGRAPH: An Open-Source Graphic Interface for Dialog Flow Design (2023.acl-demo)
Copied to clipboard
| Challenge: | Dialog systems have gained attention as a convenient way for users to access information in a more personalized manner. |
| Approach: | They present a graphical dialog flow editor built on ADVISER toolkit . it provides a clean and intuitive graphical interface for creating dialog systems . |
| Outcome: | The tool is based on the ADVISER toolkit and is evaluated with subject-experts . it is able to quickly prototype dialog systems and provide a test bed for students learning about dialog systems. |
Towards a Zero-Data, Controllable, Adaptive Dialog System (2024.lrec-main)
Copied to clipboard
| Challenge: | Recent approaches to controllable dialog systems require additional training data to be deployed in new domains. |
| Approach: | They propose to generate dialog tree data directly from dialog trees by using a commercial Large Language Model or a single GPU. |
| Outcome: | The proposed approach can achieve comparable dialog success to models trained on human data. |